BMJ Open Quality
● BMJ
Preprints posted in the last 7 days, ranked by how well they match BMJ Open Quality's content profile, based on 17 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
McHenry, R. D.; Caesar, D.; Clarke, B.; Mackay, D.; Pell, J.
Show abstract
Objectives Emergency department (ED) crowding is recognised as an important public health concern internationally, and is driven principally by exit block, the shortage of inpatient beds for patients requiring admission. This study aimed to evaluate whether a complex intervention targeting hospital occupancy improved ED patient flow, and quantified the change in attendances. Methods A controlled interrupted time series using weekly, publicly reported Public Health Scotland data from 1 January 2022 to 1 February 2026. The multi-component intervention focused on reducing hospital occupancy and included additional adult social care funding; engagement with regional social care providers; accelerated implementation of the Discharge without Delay programme; re-evaluation of whole-hospital escalation thresholds and response; resource and data supporting inpatient department reductions in length of stay; and additional investment in remote clinical assessment. The intervention commenced at a large tertiary ED on 01 February 2025. Primary outcomes were the proportions of attendances spending [≥]4, [≥]8 and [≥]12 hours in the ED. The secondary outcome was attendance volume. Segmented regression was fitted with a contemporaneous control series, seasonal terms and autoregressive moving average errors. Long waits were additionally illustrated as potentially avoided deaths. Results The analysis covered 161 pre-intervention and 52 post-intervention weeks. Relative to pre-intervention levels, the proportion of attendances waiting over 4 hours fell by 10.4% (95% CI 1.6 to 19.2%), by 16.4% (95%CI 1.3 to 31.5%) over 8 hours and by 24.3% (95%CI 2.6 to 46.1%) over 12 hours. Using established associations between long ED waits and excess mortality, by one-year the intervention was potentially associated with 54 fewer excess deaths (95%CI 19 to 93). Attendances rose by 3.8% (95%CI 1.3 to 6.4%) against the counterfactual. Conclusions A complex intervention targeting hospital occupancy was associated with a reduction in long ED waits despite rising attendances. Interventions addressing hospital occupancy can meaningfully improve ED crowding.
Jafree, D. J.; Sun, M.; Stewart, G. W.; Gishen, F.; Swanton, C.; Motallebzadeh, R.; UCL MB-PhD Outcomes Study Group,
Show abstract
Background: Clinician-scientists translate clinical observation into discovery, trials, and policy, yet this workforce is shrinking across health systems worldwide. Integrated MB-PhD training, pausing medical training to complete a PhD before clinical exposure or specialisation, is one route into this career. We aimed to evaluate the long-term value of MB-PhD training and the barriers to clinical-academic careers these face after graduation. Methods: We evaluated all 131 graduates (29.8% female) who entered the University College London (UCL) MB-PhD programme over a 25-year period (1994-2018). Bibliometric outputs were collated via an inter-linked information system. Concurrently, all 131 graduates were invited to respond to open-ended questions on career benefits and structural barriers; 99 (75.6%) responded, and responses were independently coded into themes, which were then reviewed and confirmed by a Study Group of 107 individuals, including the 91 respondents who agreed to participate further. Results: Graduates produced 5,877 publications (1,141 first-author, 819 corresponding-author), attracting 350,754 citations, with a mean relative citation ratio of 3.30 {+/-} 0.47, approximately three times the field average and sustained across three decades of programme entry. Graduates secured an estimated $157.55 million across 99 grants, released 465 public datasets, and were named investigators on 31 clinical trials across five continents. Among the 99 survey respondents, 49.5% held consultant-grade posts, 72.7% remained research-active, and 25.3% had reached senior academic grade. Open-ended responses were coded into five recurring structural barriers, subsequently confirmed by the Study Group: insufficient protected research time (72.2% of responses), unsupportive training structures and limited career opportunities (36.7%, 24.4% of responses), funding and pay barriers (22.2% of responses), and lack of mentorship or geographical/family constraints (14.4%, 13.3% of responses). Conclusions: Integrated MB-PhD training generates sustained academic productivity and leadership, but structural barriers threaten retention of graduates within clinical-academic careers. Protecting research time, stabilising funding and pay, and reducing geographic instability are needed to retain the clinician-scientists that health systems have already invested in training.
Witham, M.; Evison, F.; Bellass, S.; Cooper, R.; Gallier, S.; Pretorius, S.; Sapey, E.; Suklan, J.; Sayer, A. A.
Show abstract
Study Objective Little is known about where in hospital care for multiple long-term conditions (MLTC) is delivered. We aimed to describe pathways of care (ward transfers) and outcomes for people admitted to hospital for unscheduled care by MLTC status and other key sociodemographic characteristics. Design and setting Analysis of routinely-collected electronic health records from a large acute UK hospital. Participants Adult unscheduled care admissions from 1st July 2018 to 30th June 2019. The presence of two or more of 59 long-term conditions was ascertained using ICD-10 codes from previous hospital discharges. Main outcome measures Markov state transition probabilities were derived for ward moves and compared for MLTC vs no MLTC, age, sex, ethnicity and neighbourhood deprivation. Outcomes (length of stay, death, readmission, move from definitive ward) and time spent in emergency and assessment departments were compared between subgroups. Results A total of 33,252 adults, mean age 56.0 (SD 21.9) years were analysed; 14,834 (42.4%) had MLTC. People with MLTC were more likely to die in hospital (4.2 vs 1.9%, p<0.001), transfer to internal medicine wards or older peoples medicine wards, were less likely to transfer to surgical wards, had longer median length of stay (1.83 vs 0.69 days, p<0.001), stayed longer in acute medical units (15.5 vs 9.6 hours, p<0.001), and were more likely to move from their definitive ward (18.2 vs 16.4%, p=0.002). Conclusion Unscheduled hospital care pathways are complex and differ for people with MLTC, who have worse outcomes and may be less likely to receive optimal care.
Chowdhury, A. R.; Chowdhury, B.
Show abstract
Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.
Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.
Show abstract
Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.
Banda, M. D.; Malambo, M.
Show abstract
Despite Malawi's progress toward the UNAIDS 95-95-95 targets, facility-level rights-based challenges in HIV services persist, including stigma, discrimination and limited community participation. Community-led monitoring (CLM) has been promoted as an accountability mechanism, yet independent, facility-level evidence from urban settings remains scarce. This convergent parallel mixed-methods study assessed CLM at Ndirande and Limbe health facilities in Blantyre using a client survey (n=250), key informant interviews (n=12), and focus group discussions (three groups, 15 participants), totalling 277 participants. Chi-square tests (with Cramer's V) and binary logistic regression were used for the quantitative data; qualitative data were thematically analysed and triangulated. Analysis was guided by the rights-based approach to health and Arnstein's ladder of citizen participation. Awareness of CLM was moderate (56.0%) but participation was lower (40.2%), with involvement rated 2.78 out of 5, indicating consultative engagement. Awareness of CLM was the strongest and only robust predictor of participation (adjusted odds ratio {approx} 5.0, 95% confidence interval 2.4-10.6, p<0.001); a bivariate gender association did not survive adjustment. Notably, 41% of participants engaged in monitoring without recognising the term "CLM." CLM strengthened community-provider communication (68.5%) more than responsiveness (36.2%). Accountability mechanisms existed but functioned informally and were inconsistently documented. The two facilities did not differ significantly on any of nine indicators (all p>0.12). Barriers were structural: funding, transport, staff attitudes, fear of reprisal, and cultural norms. Urban CLM is a real but under-institutionalised accountability practice. The decisive lever is closing the awareness-action gap and formalising existing, unrecognised community monitoring through low-cost documentation, scheduled feedback, and independent, confidential complaint mechanisms. Findings are analytically transferable and offered as hypotheses for national piloting rather than as statistically generalisable conclusions.
Sahputri, V.; Angeline, A.; Tenggono, E.
Show abstract
Perioperative safety checklists standardize critical actions, but reliable completion depends on the surrounding work system and team behavior. We conducted a prospective observational analytic study from April to May 2026 in the central surgical unit of a high-volume public teaching referral hospital in Indonesia to examine whether patient safety culture and teamwork were associated with directly observed perioperative safety compliance and whether teamwork mediated the culture-compliance relationship. Patient safety culture was measured with the Hospital Survey on Patient Safety Culture 2.0, teamwork with a 35-item TeamSTEPPS Teamwork Perceptions Questionnaire research adaptation, and compliance by direct role-based observation using a 45-item checklist derived from the AORN Comprehensive Surgical Checklist. Eighty of 92 recruited professionals contributed 240 person-operation observations across 50 operations. Overall compliance was 74.75%, with sign-out lowest at 70.68%. Patient safety culture was associated with teamwork ({beta} = 0.590; 95% CI 0.510-0.770) and directly with compliance ({beta} = 0.407; 95% CI 0.187-0.712). The teamwork-compliance coefficient was positive ({beta} = 0.285; p = 0.046), but the prespecified percentile 95% CI included zero (-0.045 to 0.517). The indirect effect through teamwork was not supported ({beta} = 0.168; p = 0.079). These findings support a system-level interpretation of perioperative safety and identify learning-oriented responses to error, situation monitoring, and sign-out fidelity as measurable targets for future improvement efforts.
MURHABAZI BASHOMBWA, A.; TCHIO-NIGHIE, K. H.; NANA DJAPOU, M. C.; BUH NKUM, C.; BLAMA ABBA, I.; BEKOLO, C. E.; ATEUDJIEU, J.
Show abstract
Health facilities (HFs) routinely administer medicines and are expected to ensure patient safety by detecting, reporting, investigating, and analysing adverse events following exposure to drugs (AEFED). This study aimed to assess the implementation of pharmacovigilance activities in referral and regional health facilities in Cameroon and to identify pharmacovigilance training needs among healthcare personnel (HP). This was a cross-sectional descriptive study targeting referral and regional health facilities and healthcare personnel involved in patient care and pharmacovigilance activities in Cameroon. Health facilities were selected using stratified purposive sampling, while healthcare personnel were selected through exhaustive sampling. Data were collected using semi-structured electronic questionnaires administered face-to-face by trained enumerators. The questionnaires assessed the organization, resources, and implementation of pharmacovigilance activities at health facilities, as well as healthcare personnel knowledge of pharmacovigilance concepts, previous training, and perceived training needs. Of the 14 eligible health facilities, 10 (71.4%) consented to participate in the study. Of the 10 health facilities, 4 (40.0%) had an established pharmacovigilance unit, while 3 (30.0%) reported conducting neither detection nor notification activities. Among the 261 healthcare personnel approached, 214 (81.9%) participated. Only 41.6% had needed knowledge to detect an adverse event, while 72.9% were aware of adverse event notification procedures. Previous exposure to pharmacovigilance training was reported by 37.9% of healthcare personnel, and all participants expressed a need for additional training, particularly on national pharmacovigilance regulations (69.2%), organization of the pharmacovigilance system (67.3%), and adverse event detection (67.3%). The main reported challenges by healthcare personnel in the implementation of pharmacovigilance activities included insufficient budget allocation, limited access to pharmacovigilance training, lack of pharmacovigilance guidelines and insufficient qualified human resources. Pharmacovigilance implementation in referral and regional health facilities in Cameroon remains limited, with gaps in organizational structures, resources, healthcare personnel knowledge, and training. Strengthening pharmacovigilance systems through improved facility capacity, availability of essential tools, and targeted healthcare personnel training is needed to enhance drug safety surveillance.
McHenry, R. D.; Moultrie, C. E.
Show abstract
Objectives Emergency Department (ED) crowding is an international concern, predominantly caused by 'exit block', the lack of availability of inpatient beds for those requiring admission. The implementation of Flow Navigation Centre Plus (FNC+) services in Scotland aimed to reduce self-presentation to EDs and reduce crowding by providing remote clinical assessment for patients contacting urgent care by telephone and professional-to-professional advice on patient pathways, but their effectiveness is unknown. This study aimed to estimate the effect of board-wide implementation of FNC+ on ED attendances and long waits during the first year of FNC+ operation. Methods Controlled interrupted time series using weekly, publicly reported Public Health Scotland data. The intervention was implementation of the FNC+ in NHS Lanarkshire on 1 April 2024. Counts were summed across constituent sites and percentages derived from board totals. Co-primary outcomes were ED attendance volume and the proportions of attendances spending more than 4, 8 and 12 hours in the department. Segmented regression was fitted with contemporaneous control boards, seasonal terms, and accounted for autoregression. Results 118 pre-intervention and 52 post-intervention weeks were analysed across all 3 EDs in the implementing board. Attendances showed no detectable step change (+1.20%; 95%CIs -0.66 to +3.10) relative to the counterfactual. The estimated effect increased across follow-up, however, changing by +3.95% over 52 weeks (95% CI +0.36 to +7.67%). There was no significant step change in the proportion of attendances waiting more than 4 hours following the intervention (+1.74%; 95%CIs -0.71 to 4.20%). Some transition and structural sensitivity analyses demonstrated significant deteriorations in ED performance, and increased attendances, in the year following implementation, and none demonstrated improvements. Conclusions Board-wide implementation of a Flow Navigation Centre Plus was not associated with a step change in ED attendances or in long waits, but there is some evidence that attendances increased and long waits increased in the year following implementation. Their provision of supply-sensitive care is a possible mechanism. Additionally, given their action at the point of input, aiming to divert patients from ED attendance, it is unlikely that such services could relieve a constraint due to exit block, the availability of inpatient care for those requiring admission.
Misha, B.; Dassie, G. A.; Mohammad, I.
Show abstract
Background: Early trophic feeding promotes gut maturation, feeding tolerance, and growth in preterm neonates. However, delays remain common despite recommendations for initiation within 24 hours of birth, especially in resource-limited settings. Evidence on feeding initiation timing and predictors among Ethiopian preterm neonates is limited. Objective: To determine time to trophic feeding initiation and identify predictors among preterm neonates admitted to Adama Hospital Medical College, Ethiopia. Methods: A hospital-based retrospective cohort study was performed on 436 randomly chosen preterm neonates admitted to NICU. Data extraction was performed using a structured checklist. Time to trophic feeding initiation was analyzed using Kaplan-Meier estimates, log-rank tests, and bivariable and multivariable Cox regression models . Adjusted hazard ratios with 95% CIs were reported. Results:The sample comprised 416 preterm neonates, of whom 311 (74.8%) started trophic feeding during follow-up, and 105 (25.2%) were censored. The rate of initiation of trophic feeding was 1.92 per 100 person-hours (95% CI 1.72 to 2.15). Median time to initiation was 42 hours (interquartile range 24 to 50). Independent predictors of feeding initiation were determined by multivariable analysis and included gestational age, birth weight, maternal anaemia, respiratory distress syndrome and necrotising enterocolitis. Neonates born at 34-36 weeks had earlier initiation than those born at <34 weeks (AHR 1.39; 95 % CI 1.09 to 1.78). Similarly, neonates with a birth weight of [≥]1500 g had an earlier initiation than those with a birth weight of <1500 g (AHR 1.41; 95% CI 1.04 to 1.91). Delayed initiation was associated with maternal anaemia (AHR 0.70; 95% CI 0.51-0.95), respiratory distress syndrome (AHR 0.67; 95% CI 0.51-0.88) and necrotising enterocolitis (AHR 0.48; 95% CI 0.33-0.69). Conclusions: Delayed trophic feeding remains common among preterm neonates. Standardized feeding protocols, strengthened maternal care, and individualized nutrition strategies are needed to improve neonatal outcomes in study area.
Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.
Show abstract
The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.
Shen, H.; Agorinya, I. A.; Ayanore, M. A.; Brede, M.; Chapman, A.; Head, M.
Show abstract
Introduction Safe and timely blood availability remains a major global health challenge, especially in low- and middle-income countries. Digital tools may accelerate donor contact, but digital reachability alone does not ensure that people will notice, trust and act on urgent requests to support blood donation efforts. We examined factors associated with anticipated engagement in digitally coordinated urgent blood-donor mobilisation among digitally reachable adults in Ghana. Methods We conducted a cross-sectional online survey from September 2025 to January 2026 across Ghana's 16 regions. Participants were recruited via Facebook advertising and snowball sampling. Factors associated with urgent blood-donor mobilisability were assessed under four criteria: high future-donation willingness; high willingness to install a trusted donation app; high willingness to respond to a trusted urgent-request; and high practical flexibility to leave current activities. Descriptive analyses and multivariable logistic regression examined prevalence and associated factors. Results Among 1,067 participants, 577 (54.1%) met all four criteria. Future-donation willingness (91.8%), trusted-app installation willingness (83.2%) and trusted-request response willingness (82.7%) were common, whereas practical flexibility was lower (66.6%). In the adjusted model, high formal health-system trust (adjusted OR (AOR) 3.95, 95% CI 2.08-7.50), high digital-response readiness (AOR 2.26, 1.66-3.08), previous donation (AOR 1.47, 1.08-2.01), high donation knowledge (AOR 1.42, 1.03-1.97) and willingness to donate to strangers were positively associated with high mobilisability. Women (AOR 0.60, 0.43-0.83), participants reporting a work-schedule barrier (AOR 0.43, 0.29-0.66) and those travelling over 30 min to the nearest healthcare facility at night (AOR 0.66, 0.45-0.96) had lower adjusted odds. Conclusions Digital reachability and stated donation willingness may overestimate the population pool available for emergency donation. Digital blood-donor solutions should consider verifiable health-system requests, account for response readiness and current availability, and connect willing individuals with accessible collection options and transport support where needed.
Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.
Show abstract
BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.
Pinedo-Torres, I.; Taype-Rondan, A.; Vera-Luza, A. A.; Zegarra-Lizana, P. A.; Rojas-Vilca, J. L.; Yovera-Aldana, M.
Show abstract
Objective. To determine the publication rate of abstracts presented at the American Diabetes Association Scientific Sessions and to evaluate the association between statistical significance of study results and subsequent publication. Research Design and Methods. We conducted a retrospective cohort study of abstracts presented at the 2018 American Diabetes Association Scientific Sessions. The primary exposure was study result category (statistically significant vs. non-statistically significant findings), and the primary outcome was publication in an indexed journal within 5 years after conference presentation. Publication status was determined through PubMed/MEDLINE and Scopus searches. Adjusted relative risks (RRs) and 95% CIs were estimated using generalized linear models with Poisson distribution and robust variance. Results. Among 541 included abstracts, 321 (59.3%) were subsequently published in indexed journals. Abstracts reporting statistically significant findings had a higher publication rate than those reporting non-statistically significant findings (61.9% vs. 42.3%; p=0.002). In the adjusted analysis, abstracts with non-statistically significant findings had a lower likelihood of publication compared with those reporting statistically significant findings (adjusted RR 0.71 [95% CI 0.55-0.93]; p=0.013). Conclusions. Approximately four in ten abstracts presented at the ADA Scientific Sessions were not published within 5 years. Abstracts reporting non-statistically significant findings had a lower likelihood of subsequent publication, suggesting persistent publication bias in diabetology research. Future initiatives promoting the interpretation of effect estimates, confidence intervals and clinical relevance, rather than statistical significance alone, may help reduce selective dissemination of evidence
Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.
Show abstract
Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [≥]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.
Jawhara, B.; Baatiema, L.
Show abstract
Background: Cancer is a growing public health challenge in Ghana, with 27,385 new cases and 17,944 deaths recorded in 2022. Ghana developed a National Cancer Control Strategy (NCCS) in 2011 to guide prevention, early detection, treatment, and palliative care. The strategy expired in 2016 and has not been formally evaluated or renewed, leaving cancer control efforts without a guiding policy framework for nearly a decade. This study examined how the strategy was implemented, what barriers were encountered and what stakeholders recommend for a strengthened national cancer response. Methods: We conducted a qualitative descriptive study using semi-structured key informant interviews. Fifteen participants were recruited through purposive sampling, supplemented by snowball referrals, representing three groups: Ministry of Health policymakers, frontline healthcare providers and representatives of cancer-focused non-governmental organisations. Data were collected between June and September 2025 and analysed using Braun and Clarke's six-phase thematic analysis framework, guided deductively by the WHO Health Systems Building Blocks framework Results: Three themes emerged: NCCS interventions and systems implemented, capturing progress in cancer awareness, HPV vaccination and pilot screening programmes alongside persistent geographic and financial inequities in access; barriers to implementation, including inadequate financing, infrastructure and workforce shortages, the absence of a national cancer registry and governance failures, among them the finding that no frontline healthcare provider interviewed had any awareness of the NCCS; and recommended implementation strategies, including co-production of a renewed strategy, establishment of a dedicated National Cancer Control Programme, expanded health insurance coverage and decentralisation of oncology services. Conclusion: The NCCS was not operationally embedded in the health system. The evidence points to failures in policy dissemination as a constraint that precedes resource constraints. Addressing Ghana's rising cancer burden requires renewed political commitment, co-produced governance structures and accountability mechanisms. These findings have relevance for other low- and middle-income country settings facing similar challenges.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Manikam, L.; Fatima, A.; Patil, P.; Mayadewi, C. A.; El Khatib, T.; Drazdzewska, J.; Oyebode, O.; Llewellyn, C. H.; Webb-Martin, K.; Irish, C.; Archibong, M.; Gilmour, J.; Kalungi, P.; Batura, N.; Shringarpure, K.; Lakhanpaul, M.; Heys, M.; NEON Steering Team,
Show abstract
South Asian communities in the UK experience disproportionate maternal and child health inequalities linked to non-recommended infant feeding practices, limited health literacy, and socioeconomic constraints. Participatory learning and action (PLA) is effective in low- and middle-income countries, but high-income evidence is scarce. This pilot assessed the feasibility of a community facilitator-led PLA intervention to improve infant feeding among South Asian families in East London. A three-arm pilot feasibility cluster randomised controlled trial (ISRCTN10234623) was conducted in Tower Hamlets and Newham, East London (May-September 2022), with 12 wards randomised 1:1:1 to face-to-face PLA, online PLA, or usual care. Multilingual community facilitators delivered eight biweekly sessions over 14 weeks. Feasibility outcomes were assessed against prespecified Go/Stop criteria; exploratory outcomes included child feeding behaviours (Children's Eating Behaviour Questionnaire, CEBQ), parental feeding style (Parental Feeding Style Questionnaire, PFSQ), and child BMI Z-scores. Of 263 enrolled participants, 261 had a recorded trial arm allocation; consent to the pilot feasibility study was 70.7% (186/263; 95% CI 65.0-75.9%) meeting the [≥]50% Go criterion. Attendance was 37% (Tower Hamlets 59%, Newham 29%), below the [≥]80% Go threshold. Six-month retention was 54.8% (Tower Hamlets 78%, Newham 48.5%; 95% CI 41.8-55.3%), triggering the Definite Stop criterion. Significant baseline imbalances included BMI Z-score (p = 0.005), ethnicity, borough, and education; no between-arm BMI differences were observed at follow-up (p = 0.249). CEBQ and PFSQ baseline completion was 24.5% and 23.0%, with no usable follow-up data. PLA Phases 3 and 4 were not completed by any group; all participants providing feedback reported it acceptable. Recruitment was feasible and the intervention acceptable, but a Definite Stop criterion was triggered in Newham, no group completed the full PLA cycle, and outcome data were insufficient for evaluation. A definitive trial requires stratified randomisation, digitised multilingual data collection, participant reimbursement, and explicit PLA phase-completion criteria.
Mwenda, R. B.; Seif, S. A.; Stephano, R. O.; Moshi, F. V.
Show abstract
Background Birth Preparedness and Complication Readiness (BPCR) education is an important component of antenatal care. However, health education materials translated from one language into another may lose their intended meaning if linguistic, cultural, experiential, and sociolinguistic differences are not considered. In Tanzania, maternal health education is commonly delivered in Swahili, while many source materials are developed in English. This study explored the cultural and linguistic equivalence of BPCR terminology and concepts in a translated Swahili BPCR education manual among Hehe pregnant women in the rural Iringa Region, Tanzania. Methods A descriptive qualitative study was conducted in seven villages across Kilolo and Mufindi districts of Iringa Region. Seven focus group discussions (FGDs) involving 56 pregnant women were conducted. Participants were purposively selected from the Hehe community and were asked to interpret terminology and concepts contained in a harmonized Swahili BPCR education manual. The translation and adaptation process comprised six sequential steps: forward translation, synthesis, back translation, expert review, community exploration, and finalization. FGDs were conducted in Swahili by trained facilitators fluent in both Swahili and Hehe, audio-recorded with consent, transcribed, and thematically analyzed using Braun and Clarkes six-phase approach. Analysis focused on semantic, conceptual, experiential, and sociolinguistic equivalence. Reporting was informed by the Consolidated Criteria for Reporting Qualitative Research (COREQ). Results Five themes were developed: (1) culturally and linguistically familiar expressions conveyed BPCR concepts; (2) experiential and contextual language shaped descriptions of danger signs; (3) sociolinguistic norms and modesty influenced communication about sensitive health topics; (4) some clinically important concepts had partial or limited conceptual equivalence; and (5) unfamiliar concepts required supplementary explanation. Participants identified culturally familiar expressions including "matazamio ya kujifungua" ("anticipated date of delivery"), "fedha ndiyo usafiri" ("money itself is transport"), "chupa imepasuka" ("the water bag has burst"), "mtoto kutokucheza tumboni" ("the baby is not moving in the womb"), and "sehemu za siri" ("private parts"). Some expressions were familiar but broader than their biomedical equivalents, while cord prolapse and neonatal cyanosis had no readily recognized community equivalents. Conclusion The findings indicate that cultural and linguistic equivalence cannot be achieved through literal translation alone. Community exploration identified expressions that were familiar and socially acceptable while also revealing clinical concepts requiring additional explanation. The findings informed refinement of the Swahili BPCR education manual while preserving the intended clinical meaning. The adapted terminology should subsequently be evaluated separately for its effects on knowledge, attitudes, practices, and other health outcomes.
Yao, R.; Wi, C.-I.; Beenken, M. J.; Watson, D.; Wheeler, P. H.; Finch, M.; Kelleher, D. P.; Anil, G.; Anderson, T.; Madden, K.; Okuno, S. H.; Odedina, F. T.; Westfall, E. C.; Park, E. Y.; Sharma, P.; Dugani, S.; Foss, R. M.; Hidaka, B. H.; Sosso, J. L.; Sabarish, S.; Singh, G.; Lugo-Fagundo, N.; Howick, J.; Kim, W. R.; Calvin, A. D.; Walker-Mcgill, C. L.; Rennert, L.; Juhn, Y. J.; Cerhan, J. R.; Lynch, B. A.
Show abstract
Purpose: This study assesses the association between colorectal cancer (CRC) screening and a validated, housing-based measure of individual-level socioeconomic status (SES, called HOUSES hereafter) within rural communities and determines whether HOUSES-integrated geospatial analysis can be used to tailor interventions. Methods: We used CRC screening data from a subset of Mayo Clinic Midwest patients living in cities without ready access to routine care in the Mayo Clinic Health System in 2019 to represent rural communities. At the individual level, we assessed the association between CRC screening rates and the HOUSES index, adjusting for age, sex, race/ethnicity, comorbidity, distance from home address to clinic, and area deprivation index, using a multilevel mixed-effects logistic regression model. Additionally, we conducted geospatial analysis to examine the correlation between hotspots of 1) lower CRC screening rates and 2) lower SES of the subject population (HOUSES quartile 1). Findings: Among 34,489 individuals (median age 64.0 years, 52.4% female), those with the lowest SES (HOUSES Q1) had 37% lower odds of being CRC screening adherent than those with the highest SES (HOUSES Q4) (adj. OR [95% CI]: 0.63 [0.58-0.69]). In the 14 identified HOUSES Q1 hotspots, there was a significant correlation in counts of HOUSES Q1 and low CRC screening (correlation coefficient=0.81). Conclusion: Lower SES was significantly associated with lower CRC screening among rural populations. HOUSES-enabled geospatial analysis identified geographic hotspots with lower CRC screening rates for targeted interventions to address disparities in CRC screening in rural communities. HOUSES may be a useful digital tool for cancer preventive care and research.